Papers with safety metrics
Evaluating Psychological Safety of Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | a recent study evaluated the psychological safety of large language models. |
| Approach: | They designed unbiased prompts to evaluate the psychological safety of large language models. |
| Outcome: | The proposed prompts showed that they were fine-tuned with behavioral metrics to reduce toxicity. |
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment (2024.emnlp-main)
Copied to clipboard
Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu
| Challenge: | Large vision-language models (LVLMs) are evolving rapidly and require data with human supervision to achieve better alignment. |
| Approach: | They introduce VLFeedback, the first large-scale vision-language feedback dataset . they train Silkie, an LVLM fine-tuned via direct preference optimization . |
| Outcome: | The proposed model outperforms its base model in helpfulness, visual faithfulness, and safety metrics and exhibits enhanced resilience against red-teaming attacks. |
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Soteria locates and minimally adjusts the “functional heads” most responsible for harmful content generation in each language. |
| Approach: | Soteria locates and minimally adjusts the "functional heads" responsible for harmful content generation in each language. |
| Outcome: | The proposed approach reduces harmful content generation in languages while preserving model performance. |
The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) demonstrate their utility in character simulations, but they pose a risk of generating unsafe content. |
| Approach: | They propose a method which dynamically adjusts safety-utility preferences based on the degree of risk coupling and guides the model to generate responses biased toward utility or safety. |
| Outcome: | The proposed method improves safety metrics while maintaining utility. |
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) face challenges in maintaining accuracy due to the dynamic nature of world knowledge. |
| Approach: | They propose to use a benchmark dataset to investigate the effects of model edits on model safety metrics and guardrails. |
| Outcome: | The proposed dataset sheds light on how the edits, impact the model’s safety metrics and guardrails. |
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing safety evaluations rely on artificial images to evaluate vision-language models . a recent study found that memes are more effective at bypassing safety measures than synthetic or typographic images. |
| Approach: | They propose a benchmark pairing meme images with harmful and benign instructions . they assess multiple VLMs across single and multi-turn interactions . |
| Outcome: | The proposed benchmark pairs real meme images with harmful and benign instructions. |